Micron Document
<!DOCTYPE html>
<html class="client-nojs vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-0 vector-toc-not-available vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-0 skin-theme-clientpref-day vector-sticky-header-enabled" lang="de" dir="ltr"><head>
<meta charset="UTF-8">
<title>Convolutional Neural Network</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="icon" type="image/png" href="./_res_/favicon.png">
<link rel="canonical" href="https://de.wikipedia.org/wiki/Convolutional_Neural_Network"> <link href="./_mw_/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.math.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.wikimediamessages.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./_mw_/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./_mw_/skins.vector.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link href="./_mw_/ext.gadget.citeRef.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.defaultPlainlinks.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiCommonHide.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiCommonLayout.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiCommonStyle.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiDarkmode.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.dewikiResponsive.css" rel="stylesheet" type="text/css">
<link href="./_mw_/ext.gadget.specialSearch.css" rel="stylesheet" type="text/css">
<link rel="stylesheet" type="text/css" href="./_mw_/site.styles.css">
<link rel="stylesheet" type="text/css" href="./_mw_/noscript.css">
<link rel="stylesheet" type="text/css" href="./_res_/footer.css">
<link rel="stylesheet" type="text/css" href="./_res_/vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Convolutional_Neural_Network rootpage-Convolutional_Neural_Network skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading"><span class="mw-page-title-main">Convolutional Neural Network</span></h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="contentSub">
<div id="mw-content-subtitle"></div>
</div>
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="de" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="de" dir="ltr"><p>Ein <b>Convolutional Neural Network</b> (<b>CNN</b> oder <b>ConvNet</b>), zu Deutsch etwa „<a href="Faltung_(Mathematik)" title="Faltung (Mathematik)">faltendes</a> neuronales Netzwerk“, ist ein <a href="K%C3%BCnstliches_neuronales_Netz" title="Künstliches neuronales Netz">künstliches neuronales Netz</a>. Es handelt sich um ein von biologischen Prozessen inspiriertes Konzept im Bereich des <a href="Maschinelles_Lernen" title="Maschinelles Lernen">maschinellen Lernens</a><sup id="cite_ref-robust_face_detection_1-0" class="reference"><a href="#cite_note-robust_face_detection-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>. <i>Convolutional Neural Networks</i> finden Anwendung in zahlreichen Technologien der <a href="K%C3%BCnstliche_Intelligenz" title="Künstliche Intelligenz">künstlichen Intelligenz</a>, vornehmlich bei der maschinellen Verarbeitung von Bild- oder Audiodaten.
</p><p>Die CNN-Architektur wurde 1979 von <a href="Kunihiko_Fukushima" title="Kunihiko Fukushima">Kunihiko Fukushima</a> unter dem Namen <a href="Neocognitron" title="Neocognitron">Neocognitron</a> eingeführt.<sup id="cite_ref-fukushima1980_2-0" class="reference"><a href="#cite_note-fukushima1980-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-schmidhuber2015_3-0" class="reference"><a href="#cite_note-schmidhuber2015-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> Im Jahr 1987 trainierte <a href="Alex_Waibel" class="mw-redirect" title="Alex Waibel">Alex Waibels</a> ein CNN namens TDNN durch <a href="Backpropagation" title="Backpropagation">Backpropagation</a> und erzielte damit Bewegungsinvarianz.<sup id="cite_ref-Waibel1987_4-0" class="reference"><a href="#cite_note-Waibel1987-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> Auch <a href="Yann_LeCun" title="Yann LeCun">Yann LeCun</a> publizierte ab dem Ende der 1980er Jahre wichtige Beiträge zu CNNs.<sup id="cite_ref-lecun1989_5-0" class="reference"><a href="#cite_note-lecun1989-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup>
</p>

<div class="mw-heading mw-heading2"><h2 id="Aufbau">Aufbau</h2></div>

<p>Grundsätzlich besteht die Struktur eines klassischen Convolutional Neural Networks aus einem oder mehreren Convolutional Layer, gefolgt von einem Pooling Layer.
Diese Einheit kann sich prinzipiell beliebig oft wiederholen, bei ausreichend Wiederholungen spricht man dann von Deep Convolutional Neural Networks, die in den Bereich <a href="Deep_Learning" title="Deep Learning">Deep Learning</a> fallen.
Hierbei ist auf die Ähnlichkeit zum <a href="Optimalfilter" title="Optimalfilter">Optimalfilter</a> hinzuweisen<sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup>.
Architektonisch können im Vergleich zum mehrlagigen Perzeptron (<a href="Multi-Layer-Perzeptron" title="Multi-Layer-Perzeptron">Multi-Layer-Perzeptron</a>) drei wesentliche Unterschiede festgehalten werden (Details hierzu siehe <i>Convolutional Layer</i>):
</p>
<ul><li>2D- oder 3D-Anordnung der Neuronen</li>
<li>Geteilte Gewichte</li>
<li>Lokale Konnektivität</li></ul>
<div class="mw-heading mw-heading3"><h3 id="Convolutional_Layer">Convolutional Layer</h3></div>
<p>In der Regel liegt die Eingabe als zwei- oder dreidimensionale Matrix (z.&nbsp;B. die <a href="Pixel" title="Pixel">Pixel</a> eines Graustufen- oder Farbbildes) vor. Dementsprechend sind die Neuronen im Convolutional Layer angeordnet.
</p><p>Die Aktivität jedes Neurons wird über eine diskrete <a href="Faltung_(Mathematik)" title="Faltung (Mathematik)">Faltung</a> (daher der Zusatz <i>convolutional</i>) berechnet. Dabei wird schrittweise eine vergleichsweise kleine <a href="Faltungsmatrix" title="Faltungsmatrix">Faltungsmatrix</a> (Filterkernel) über die Eingabe bewegt. Die Eingabe eines Neurons im Convolutional Layer berechnet sich als <a href="Skalarprodukt" title="Skalarprodukt">inneres Produkt</a> des Filterkernels mit dem aktuell unterliegenden Bildausschnitt. Dementsprechend reagieren benachbarte Neuronen im Convolutional Layer auf sich überlappende Bereiche (ähnliche Frequenzen in Audiosignalen oder lokale Umgebungen in Bildern).<sup id="cite_ref-deeplearning_8-0" class="reference"><a href="#cite_note-deeplearning-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup>
</p>

<p>Hervorzuheben ist, dass ein Neuron in diesem Layer nur auf Reize in einer lokalen Umgebung des vorherigen Layers reagiert. Dies folgt dem biologischen Vorbild des <a href="Rezeptives_Feld" title="Rezeptives Feld">rezeptiven Feldes</a>. Zudem sind die Gewichte für alle Neuronen eines Convolutional Layers identisch (geteilte Gewichte, englisch: <i>shared weights</i>). Dies führt dazu, dass beispielsweise jedes Neuron im ersten Convolutional Layer codiert, zu welcher Intensität eine Kante in einem bestimmten lokalen Bereich der Eingabe vorliegt. Die Kantenerkennung als erster Schritt der Bilderkennung besitzt hohe biologische <a href="Plausibilit%C3%A4t" title="Plausibilität">Plausibilität</a>.<sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup> Aus den <i>shared weights</i> folgt unmittelbar, dass <a href="Translationsinvarianz" class="mw-redirect" title="Translationsinvarianz">Translationsinvarianz</a> eine inhärente Eigenschaft von CNNs ist.
</p><p>Der mittels diskreter Faltung ermittelte Input eines jeden Neurons wird nun von einer Aktivierungsfunktion, bei CNNs üblicherweise <a href="Rectifier_(neuronale_Netzwerke)" title="Rectifier (neuronale Netzwerke)">Rectified Linear Unit</a>, kurz ReLU (<span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle f(x)=\max(0,x)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>f</mi>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo movablelimits="true" form="prefix">max</mo>
<mo stretchy="false">(</mo>
<mn>0</mn>
<mo>,</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle f(x)=\max(0,x)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/5fa5d3598751091eed580bd9dca873f496a2d0ac.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:17.177ex; height:2.843ex;" alt="{\displaystyle f(x)=\max(0,x)}" loading="lazy"></span>), in den Output verwandelt, der die relative Feuerfrequenz eines echten Neurons modellieren soll. Da <a href="Backpropagation" title="Backpropagation">Backpropagation</a> die Berechnung der <a href="Gradient_(Mathematik)" title="Gradient (Mathematik)">Gradienten</a> verlangt, wird in der Praxis eine differenzierbare Approximation von ReLU benutzt: <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle f(x)=\ln(1+e^{x})}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>f</mi>
<mo stretchy="false">(</mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mi>ln</mi>
<mo>⁡<!-- ⁡ --></mo>
<mo stretchy="false">(</mo>
<mn>1</mn>
<mo>+</mo>
<msup>
<mi>e</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
</mrow>
</msup>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle f(x)=\ln(1+e^{x})}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/f21f3d1e2c67c5c2d2085e384512bc737c8e14af.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:17.524ex; height:2.843ex;" alt="{\displaystyle f(x)=\ln(1+e^{x})}" loading="lazy"></span>
</p><p>Analog zum visuellen Cortex steigt in tiefer gelegenen Convolutional Layers sowohl die Größe der rezeptiven Felder (siehe Sektion <i>Pooling Layer</i>) als auch die Komplexität der erkannten Features (beispielsweise Teile eines Gesichts).
</p><p>Eine Faltung ist die einfache Anwendung eines Filters auf eine Eingabe, die zu einer Aktivierung führt. Die wiederholte Anwendung desselben Filters auf eine Eingabe führt zu einer Karte von Aktivierungen, die als <i>Feature Map</i> bezeichnet wird und die Positionen und die Stärke eines erkannten Features in einer Eingabe, beispielsweise einem Bild, angibt. Convolutional Layer können automatisch eine große Anzahl von Filtern parallel lernen, die für einen Trainingsdatensatz spezifisch sind. Das Ergebnis sind hochspezifische Merkmale, die überall auf Eingabebildern erkannt werden können. Convolutional Layer wenden einen Filter auf eine Eingabe an, um eine Feature Map zu erstellen, die das Vorhandensein erkannter Features in der Eingabe zusammenfasst.
</p><p>Eine Faltung ist eine lineare Operation, die die Multiplikation einer Reihe von Gewichten mit der Eingabe beinhaltet, ähnlich wie bei einem herkömmlichen <a href="Neuronales_Netz" title="Neuronales Netz">neuronalen Netz</a>. Der Filter ist kleiner als die Eingabedaten und die Art der Multiplikation, die zwischen einem filtergroßen Patch der Eingabe und dem Filter angewendet wird, ist ein <a href="Skalarprodukt" title="Skalarprodukt">Skalarprodukt</a>. Die Verwendung eines Filters, der kleiner als die Eingabe ist, ist beabsichtigt, weil dadurch derselbe Filter an verschiedenen Stellen der Eingabe mehrmals mit dem Eingabearray multipliziert werden kann. Konkret wird der Filter systematisch auf jeden überlappenden Teil oder filtergroßen Patch der Eingabedaten angewendet, von links nach rechts, von oben nach unten. Diese systematische Anwendung desselben Filters auf ein Bild ist eine wirkungsvolle Idee. Wenn der Filter darauf ausgelegt ist, einen bestimmten Merkmalstyp in der Eingabe zu erkennen, kann der Filter dieses Merkmal an einer beliebigen Stelle im Bild erkennen. Diese Fähigkeit wird allgemein als Übersetzungsinvarianz bezeichnet, zum Beispiel das allgemeine Interesse daran, ob das Merkmal vorhanden ist, und nicht, wo es vorhanden war.<sup id="cite_ref-10" class="reference"><a href="#cite_note-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup>
</p><p>Die Convolutional Layer ist der Kernbaustein des Convolutional Neural Network. Es trägt den Hauptanteil der Rechenlast des Netzwerks. Diese Schicht führt ein <a href="Skalarprodukt" title="Skalarprodukt">Skalarprodukt</a> zwischen zwei Matrizen aus, wobei eine <a href="Matrix_(Mathematik)" title="Matrix (Mathematik)">Matrix</a> der Satz von gelernten Parametern ist, die auch als <i>Kernel</i> bezeichnet werden, und die andere Matrix der eingeschränkte Teil des Empfangsfeldes. Der Kernel ist räumlich kleiner als ein Bild, ist aber ausführlicher. Dies bedeutet, dass, wenn das Bild aus drei Kanälen besteht, die Höhe und die Breite des Kernels räumlich klein sind, die Tiefe bis zu allen drei Kanälen erstreckt. Dies erzeugt eine zweidimensionale Darstellung des als Aktivierungskarte bekannten Bildes, die die Antwort des Kernels an jeder räumlichen Position des Bildes ergibt. Die Gleitgröße des Kernels wird als Schritt bezeichnet.
</p><p>Wenn man eine Eingabe der Größe <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle W\times W\times D}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>W</mi>
<mo>×<!-- × --></mo>
<mi>W</mi>
<mo>×<!-- × --></mo>
<mi>D</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle W\times W\times D}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/2109ee94572f6e74bd1c9ab48836f309e4bdbd33.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:12.476ex; height:2.176ex;" alt="{\displaystyle W\times W\times D}" loading="lazy"></span> und <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle D_{out}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>D</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>o</mi>
<mi>u</mi>
<mi>t</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle D_{out}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c9020bc4fdf75458ecd909c1186fa2b93a421e60.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:4.488ex; height:2.509ex;" alt="{\displaystyle D_{out}}" loading="lazy"></span> Kerne mit einer räumlichen Größe von <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle F}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>F</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle F}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/545fd099af8541605f7ee55f08225526be88ce57.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.741ex; height:2.176ex;" alt="{\displaystyle F}" loading="lazy"></span> mit Schrittweite <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle S}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>S</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle S}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/4611d85173cd3b508e67077d4a1252c9c05abca2.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.499ex; height:2.176ex;" alt="{\displaystyle S}" loading="lazy"></span> und dem Betrag der Füllung <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle P}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>P</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle P}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/b4dc73bf40314945ff376bd363916a738548d40a.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.745ex; height:2.176ex;" alt="{\displaystyle P}" loading="lazy"></span> hat, kann die Größe des Ausgabevolumens durch die folgende Formel bestimmt werden:<sup id="cite_ref-:0_11-0" class="reference"><a href="#cite_note-:0-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup>
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle W_{out}={\frac {W-F+2\cdot P}{S}}+1}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>W</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>o</mi>
<mi>u</mi>
<mi>t</mi>
</mrow>
</msub>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mrow>
<mi>W</mi>
<mo>−<!-- − --></mo>
<mi>F</mi>
<mo>+</mo>
<mn>2</mn>
<mo>⋅<!-- ⋅ --></mo>
<mi>P</mi>
</mrow>
<mi>S</mi>
</mfrac>
</mrow>
<mo>+</mo>
<mn>1</mn>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle W_{out}={\frac {W-F+2\cdot P}{S}}+1}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/517b8b6c3b6093a9c6cddcc852bc9529f0d28098.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -2.005ex; width:27.138ex; height:5.343ex;" alt="{\displaystyle W_{out}={\frac {W-F+2\cdot P}{S}}+1}" loading="lazy"></span></dd></dl>
<div class="mw-heading mw-heading3"><h3 id="Pooling_Layer">Pooling Layer</h3></div>

<p>Im folgenden Schritt, dem Pooling, werden überflüssige Informationen verworfen. Zur Objekterkennung in Bildern etwa ist die <i>exakte</i> Position einer Kante im Bild von vernachlässigbarem Interesse – die ungefähre Lokalisierung eines Features ist hinreichend.
Es gibt verschiedene Arten des Poolings. Mit Abstand am stärksten verbreitet ist das Max-Pooling<sup id="cite_ref-weng1993_12-0" class="reference"><a href="#cite_note-weng1993-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Yamaguchi1990_13-0" class="reference"><a href="#cite_note-Yamaguchi1990-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup>, wobei aus jedem 2×2-Quadrat aus Neuronen des Convolutional Layers nur die Aktivität des aktivsten (daher „Max“) Neurons für die weiteren Berechnungsschritte beibehalten wird; die Aktivität der übrigen Neuronen wird verworfen (siehe Bild).
Trotz der Datenreduktion (im Beispiel 75&nbsp;%) verringert sich in der Regel die Performance des Netzwerks nicht durch das Pooling. Im Gegenteil, es bietet einige signifikante Vorteile:
</p>
<ul><li>Verringerter Platzbedarf und erhöhte Berechnungsgeschwindigkeit</li>
<li>Daraus resultierende Möglichkeit zur Erzeugung tieferer Netzwerke, die komplexere Aufgaben lösen können</li>
<li>Automatisches Wachstum der Größe der rezeptiven Felder in tieferen Convolutional Layers (ohne dass dafür explizit die Größe der Faltungsmatrizen erhöht werden müsste)</li>
<li>Präventionsmaßnahme gegen <a href="%C3%9Cberanpassung" title="Überanpassung">Overfitting</a></li></ul>
<p>Alternativen wie das Mean-Pooling haben sich in der Praxis als weniger effizient erwiesen.<sup id="cite_ref-Scherer_15-0" class="reference"><a href="#cite_note-Scherer-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup>
</p><p>Das biologische Pendant zum Pooling ist die <a href="Laterale_Hemmung" title="Laterale Hemmung">laterale Hemmung</a> im visuellen Cortex.
</p><p>Die Pooling Layer ersetzt die Ausgabe des Netzwerks an bestimmten Stellen, indem er eine zusammenfassende Statistik der nahe gelegenen Ausgaben abgeleitet hat. Dies hilft bei der Verringerung der räumlichen Größe der Darstellung, die die erforderliche Menge an Rechen und Gewichten verringert. Die Pooling-Operation wird auf jedem Stück der Darstellung einzeln ausgeführt. Das beliebteste Verfahren ist Max Pooling, die die maximale Ausgabe aus der Nachbarschaft angibt. Wenn man eine Aktivierungskarte der Größe <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle W\times W\times D}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>W</mi>
<mo>×<!-- × --></mo>
<mi>W</mi>
<mo>×<!-- × --></mo>
<mi>D</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle W\times W\times D}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/2109ee94572f6e74bd1c9ab48836f309e4bdbd33.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:12.476ex; height:2.176ex;" alt="{\displaystyle W\times W\times D}" loading="lazy"></span>, einen Pooling Kernel der räumlichen Größe <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle F}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>F</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle F}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/545fd099af8541605f7ee55f08225526be88ce57.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.741ex; height:2.176ex;" alt="{\displaystyle F}" loading="lazy"></span> mit Schrittweite <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle S}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>S</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle S}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/4611d85173cd3b508e67077d4a1252c9c05abca2.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.499ex; height:2.176ex;" alt="{\displaystyle S}" loading="lazy"></span> hat, kann die Größe des Ausgabevolumens durch die folgende Formel bestimmt werden:<sup id="cite_ref-:0_11-1" class="reference"><a href="#cite_note-:0-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup>
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle W_{out}={\frac {W-F}{S}}+1}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>W</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>o</mi>
<mi>u</mi>
<mi>t</mi>
</mrow>
</msub>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mrow>
<mi>W</mi>
<mo>−<!-- − --></mo>
<mi>F</mi>
</mrow>
<mi>S</mi>
</mfrac>
</mrow>
<mo>+</mo>
<mn>1</mn>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle W_{out}={\frac {W-F}{S}}+1}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/a688bfab01bf03ce7893dfb0625ffb47c43a43fc.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -2.005ex; width:19.711ex; height:5.343ex;" alt="{\displaystyle W_{out}={\frac {W-F}{S}}+1}" loading="lazy"></span></dd></dl>
<p>Dies ergibt ein Ausgabevolumen der Größe <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle W_{out}\times W_{out}\times D}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>W</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>o</mi>
<mi>u</mi>
<mi>t</mi>
</mrow>
</msub>
<mo>×<!-- × --></mo>
<msub>
<mi>W</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>o</mi>
<mi>u</mi>
<mi>t</mi>
</mrow>
</msub>
<mo>×<!-- × --></mo>
<mi>D</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle W_{out}\times W_{out}\times D}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/131115ecdeedb551a80b8968ab1847d6fdfa5727.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:17.119ex; height:2.509ex;" alt="{\displaystyle W_{out}\times W_{out}\times D}" loading="lazy"></span>.
</p>
<div class="mw-heading mw-heading4"><h4 id="Max_Pooling">Max Pooling</h4></div>
<p>Max Pooling ist eine Faltungstechnik, die den Maximalwert aus dem Patch der Eingabedaten auswählt und diese Werte in einer Feature Map zusammenfasst: Diese Methode behält die wichtigsten Merkmale der Eingabe bei, indem sie ihre Abmessungen reduziert. Die mathematische Formel für Max Pooling lautet:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \mathrm {MaxPooling} (x)_{i,j,k}=\max _{m,n}{X_{i\cdot x_{s}+m,j\cdot y_{s}+n,k}}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">M</mi>
<mi mathvariant="normal">a</mi>
<mi mathvariant="normal">x</mi>
<mi mathvariant="normal">P</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">n</mi>
<mi mathvariant="normal">g</mi>
</mrow>
<mo stretchy="false">(</mo>
<mi>x</mi>
<msub>
<mo stretchy="false">)</mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
<mo>=</mo>
<munder>
<mo movablelimits="true" form="prefix">max</mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
<mo>,</mo>
<mi>n</mi>
</mrow>
</munder>
<mrow class="MJX-TeXAtom-ORD">
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>s</mi>
</mrow>
</msub>
<mo>+</mo>
<mi>m</mi>
<mo>,</mo>
<mi>j</mi>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>y</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>s</mi>
</mrow>
</msub>
<mo>+</mo>
<mi>n</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \mathrm {MaxPooling} (x)_{i,j,k}=\max _{m,n}{X_{i\cdot x_{s}+m,j\cdot y_{s}+n,k}}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/d1a905567808d8fad296e4b517d20faf33961ed0.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -2.171ex; width:40.776ex; height:4.176ex;" alt="{\displaystyle \mathrm {MaxPooling} (x)_{i,j,k}=\max _{m,n}{X_{i\cdot x_{s}+m,j\cdot y_{s}+n,k}}}" loading="lazy"></span></dd></dl>
<p>Dabei sind <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/68baa052181f707c662844a465bfeeb135e82bab.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.98ex; height:2.176ex;" alt="{\displaystyle X}" loading="lazy"></span> die Eingabe, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle (i,j)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle (i,j)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/8ef21910f980c6fca2b15bee102a7a0d861ed712.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:4.604ex; height:2.843ex;" alt="{\displaystyle (i,j)}" loading="lazy"></span> die Indexe der Ausgabe, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> der Kanalindex, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle s_{x}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>s</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle s_{x}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/f738a5f93f694ded413a8f8da71081d0fdff4501.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:2.263ex; height:2.009ex;" alt="{\displaystyle s_{x}}" loading="lazy"></span> und <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle s_{y}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>s</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>y</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle s_{y}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/bec93eb58677e5642c86ecdb2430703a8ee57da2.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -1.005ex; width:2.14ex; height:2.343ex;" alt="{\displaystyle s_{y}}" loading="lazy"></span> die Schrittwerte in horizontaler bzw. vertikaler Richtung und das Pooling-Fenster wird durch die Filtergrößen <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle f_{x}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>f</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle f_{x}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/bb36605624ab2287dc5ec558513c625b88acfee3.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:2.312ex; height:2.509ex;" alt="{\displaystyle f_{x}}" loading="lazy"></span> und <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle f_{x}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>f</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle f_{x}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/bb36605624ab2287dc5ec558513c625b88acfee3.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:2.312ex; height:2.509ex;" alt="{\displaystyle f_{x}}" loading="lazy"></span> definiert, zentriert am Ausgabeindex <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle (i,j)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle (i,j)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/8ef21910f980c6fca2b15bee102a7a0d861ed712.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:4.604ex; height:2.843ex;" alt="{\displaystyle (i,j)}" loading="lazy"></span>.
</p>
<div class="mw-heading mw-heading4"><h4 id="Average_Pooling">Average Pooling</h4></div>
<p>Durch das Average Pooling wird der Durchschnittswert aus einem Bereich von Eingabedaten berechnet und diese Werte in einer Feature Map zusammengefasst. Diese Methode ist in Fällen vorzuziehen, in denen eine Glättung der Eingabedaten erforderlich ist, da sie dabei hilft, das Vorhandensein von Ausreißern zu identifizieren. Die mathematische Formel für Average Pooling lautet:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \mathrm {AvgPooling} (x)_{i,j,k}={\frac {1}{f_{x}\cdot f_{y}}}\cdot \sum _{m,n}X_{i\cdot x_{s}+m,j\cdot y_{s}+n,k}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">A</mi>
<mi mathvariant="normal">v</mi>
<mi mathvariant="normal">g</mi>
<mi mathvariant="normal">P</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">n</mi>
<mi mathvariant="normal">g</mi>
</mrow>
<mo stretchy="false">(</mo>
<mi>x</mi>
<msub>
<mo stretchy="false">)</mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mfrac>
<mn>1</mn>
<mrow>
<msub>
<mi>f</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
</mrow>
</msub>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>f</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>y</mi>
</mrow>
</msub>
</mrow>
</mfrac>
</mrow>
<mo>⋅<!-- ⋅ --></mo>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
<mo>,</mo>
<mi>n</mi>
</mrow>
</munder>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>x</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>s</mi>
</mrow>
</msub>
<mo>+</mo>
<mi>m</mi>
<mo>,</mo>
<mi>j</mi>
<mo>⋅<!-- ⋅ --></mo>
<msub>
<mi>y</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>s</mi>
</mrow>
</msub>
<mo>+</mo>
<mi>n</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \mathrm {AvgPooling} (x)_{i,j,k}={\frac {1}{f_{x}\cdot f_{y}}}\cdot \sum _{m,n}X_{i\cdot x_{s}+m,j\cdot y_{s}+n,k}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/e0a4c51919f3c3872d47083f6297835f4facf8ed.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.171ex; width:48.112ex; height:6.509ex;" alt="{\displaystyle \mathrm {AvgPooling} (x)_{i,j,k}={\frac {1}{f_{x}\cdot f_{y}}}\cdot \sum _{m,n}X_{i\cdot x_{s}+m,j\cdot y_{s}+n,k}}" loading="lazy"></span></dd></dl>
<div class="mw-heading mw-heading4"><h4 id="Global_Pooling">Global Pooling</h4></div>
<p>Global Pooling fasst die Werte aller Neuronen für jeden Patch der Eingabedaten in einer Feature Map zusammen, unabhängig von ihrer räumlichen Position. Diese Technik wird auch verwendet, um die Dimensionalität der Eingabe zu reduzieren und kann entweder mithilfe der maximalen oder durchschnittlichen Pooling-Operation durchgeführt werden. Die mathematische Formel für Global Pooling lautet:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \mathrm {GlobalPooling} (x)_{k}=f_{x}(X:,:,k)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">G</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">b</mi>
<mi mathvariant="normal">a</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">P</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">n</mi>
<mi mathvariant="normal">g</mi>
</mrow>
<mo stretchy="false">(</mo>
<mi>x</mi>
<msub>
<mo stretchy="false">)</mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>k</mi>
</mrow>
</msub>
<mo>=</mo>
<msub>
<mi>f</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>x</mi>
</mrow>
</msub>
<mo stretchy="false">(</mo>
<mi>X</mi>
<mo>:</mo>
<mo>,</mo>
<mo>:</mo>
<mo>,</mo>
<mi>k</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \mathrm {GlobalPooling} (x)_{k}=f_{x}(X:,:,k)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/4838524c059e1e4c72e5be1e37c67a9c4782948e.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:33.037ex; height:2.843ex;" alt="{\displaystyle \mathrm {GlobalPooling} (x)_{k}=f_{x}(X:,:,k)}" loading="lazy"></span></dd></dl>
<div class="mw-heading mw-heading4"><h4 id="Stochastic_Pooling">Stochastic Pooling</h4></div>
<p>Stochastic Pooling ist eine nicht-deterministische Pooling-Operation, die zufällige Werte mit Max Pooling kombiniert. Diese Technik trägt dazu bei, die Robustheit des Modells gegenüber kleinen Abweichungen in den Eingabedaten zu verbessern. Die mathematische Formel für Stochastic Pooling lautet:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle \mathrm {StochasticPooling} (x)_{i,j,k}={\begin{cases}{\begin{aligned}X_{i,j,k}&amp;\ \mathrm {mitWahrscheinlichkeit} \ p_{i,k}\\0&amp;\ \mathrm {mitWahrscheinlichkeit} \ 1-p_{i,k}\\\end{aligned}}\end{cases}}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">S</mi>
<mi mathvariant="normal">t</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">c</mi>
<mi mathvariant="normal">h</mi>
<mi mathvariant="normal">a</mi>
<mi mathvariant="normal">s</mi>
<mi mathvariant="normal">t</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">c</mi>
<mi mathvariant="normal">P</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">o</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">n</mi>
<mi mathvariant="normal">g</mi>
</mrow>
<mo stretchy="false">(</mo>
<mi>x</mi>
<msub>
<mo stretchy="false">)</mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
<mo>=</mo>
<mrow class="MJX-TeXAtom-ORD">
<mrow>
<mo>{</mo>
<mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false">
<mtr>
<mtd>
<mrow class="MJX-TeXAtom-ORD">
<mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true">
<mtr>
<mtd>
<msub>
<mi>X</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
</mtd>
<mtd>
<mtext>&nbsp;</mtext>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">m</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">t</mi>
<mi mathvariant="normal">W</mi>
<mi mathvariant="normal">a</mi>
<mi mathvariant="normal">h</mi>
<mi mathvariant="normal">r</mi>
<mi mathvariant="normal">s</mi>
<mi mathvariant="normal">c</mi>
<mi mathvariant="normal">h</mi>
<mi mathvariant="normal">e</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">n</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">c</mi>
<mi mathvariant="normal">h</mi>
<mi mathvariant="normal">k</mi>
<mi mathvariant="normal">e</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">t</mi>
</mrow>
<mtext>&nbsp;</mtext>
<msub>
<mi>p</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
</mtd>
</mtr>
<mtr>
<mtd>
<mn>0</mn>
</mtd>
<mtd>
<mtext>&nbsp;</mtext>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">m</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">t</mi>
<mi mathvariant="normal">W</mi>
<mi mathvariant="normal">a</mi>
<mi mathvariant="normal">h</mi>
<mi mathvariant="normal">r</mi>
<mi mathvariant="normal">s</mi>
<mi mathvariant="normal">c</mi>
<mi mathvariant="normal">h</mi>
<mi mathvariant="normal">e</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">n</mi>
<mi mathvariant="normal">l</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">c</mi>
<mi mathvariant="normal">h</mi>
<mi mathvariant="normal">k</mi>
<mi mathvariant="normal">e</mi>
<mi mathvariant="normal">i</mi>
<mi mathvariant="normal">t</mi>
</mrow>
<mtext>&nbsp;</mtext>
<mn>1</mn>
<mo>−<!-- − --></mo>
<msub>
<mi>p</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>k</mi>
</mrow>
</msub>
</mtd>
</mtr>
</mtable>
</mrow>
</mtd>
</mtr>
</mtable>
<mo fence="true" stretchy="true" symmetric="true"></mo>
</mrow>
</mrow>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle \mathrm {StochasticPooling} (x)_{i,j,k}={\begin{cases}{\begin{aligned}X_{i,j,k}&amp;\ \mathrm {mitWahrscheinlichkeit} \ p_{i,k}\\0&amp;\ \mathrm {mitWahrscheinlichkeit} \ 1-p_{i,k}\\\end{aligned}}\end{cases}}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/cd9a4983968e8eb6135d4e4426da316c5d9353f5.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -2.303ex; margin-bottom: -0.201ex; width:67.241ex; height:6.176ex;" alt="{\displaystyle \mathrm {StochasticPooling} (x)_{i,j,k}={\begin{cases}{\begin{aligned}X_{i,j,k}&amp;\ \mathrm {mitWahrscheinlichkeit} \ p_{i,k}\\0&amp;\ \mathrm {mitWahrscheinlichkeit} \ 1-p_{i,k}\\\end{aligned}}\end{cases}}}" loading="lazy"></span></dd></dl>
<p>Dabei ist <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/68baa052181f707c662844a465bfeeb135e82bab.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.98ex; height:2.176ex;" alt="{\displaystyle X}" loading="lazy"></span> die Eingabe, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle (i,j)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle (i,j)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/8ef21910f980c6fca2b15bee102a7a0d861ed712.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:4.604ex; height:2.843ex;" alt="{\displaystyle (i,j)}" loading="lazy"></span> die Indexe des Ausgabetensors, <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle k}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>k</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle k}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/c3c9a2c7b599b37105512c5d570edc034056dd40.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.211ex; height:2.176ex;" alt="{\displaystyle k}" loading="lazy"></span> der Kanalindex und <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle p_{i,j}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<msub>
<mi>p</mi>
<mrow class="MJX-TeXAtom-ORD">
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
</mrow>
</msub>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle p_{i,j}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/e82f4cb43c1f5cd53898ed7cc80dcf092373f676.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -1.005ex; margin-left: -0.089ex; width:3.193ex; height:2.343ex;" alt="{\displaystyle p_{i,j}}" loading="lazy"></span> die <a href="Wahrscheinlichkeit" title="Wahrscheinlichkeit">Wahrscheinlichkeit</a>, den Wert an Position <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle (i,j)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle (i,j)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/8ef21910f980c6fca2b15bee102a7a0d861ed712.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:4.604ex; height:2.843ex;" alt="{\displaystyle (i,j)}" loading="lazy"></span> in der Feature Map der Eingabe beizubehalten. Die Wahrscheinlichkeiten werden für jedes Pooling-Fenster zufällig generiert und sind normalerweise proportional zu den Werten im Fenster.
</p>
<div class="mw-heading mw-heading3"><h3 id="Fully-connected_Layer">Fully-connected Layer</h3></div>
<p>Nach einigen sich wiederholenden Einheiten bestehend aus Convolutional und Pooling Layer kann das Netzwerk mit einem (oder mehreren) Fully-connected Layer entsprechend der Architektur des mehrlagigen <a href="Perzeptron" title="Perzeptron">Perzeptrons</a> abschließen. Dies wird vor allem bei der <a href="Klassifizierung" title="Klassifizierung">Klassifizierung</a> angewendet. Die Anzahl der Neuronen im letzten Layer korrespondiert dann üblicherweise zu der Anzahl an (Objekt-)Klassen, die das Netz unterscheiden soll. Dieses, sehr redundante, sogenannte <a href="1-aus-n-Code" title="1-aus-n-Code">One-Hot-encoding</a> hat den Vorteil, dass keine impliziten Annahmen über Ähnlichkeiten von Klassen gemacht werden.
</p><p>Die Ausgabe der letzten Schicht des CNNs wird in der Regel durch eine <a href="Softmax-Funktion" title="Softmax-Funktion">Softmax-Funktion</a>, einer <a href="Translationsinvarianz" class="mw-redirect" title="Translationsinvarianz">translations</a>- aber nicht <a href="Skaleninvarianz" title="Skaleninvarianz">skaleninvarianten</a> Normalisierung über alle Neuronen im letzten Layer, in eine <a href="Wahrscheinlichkeitsma%C3%9F" title="Wahrscheinlichkeitsmaß">Wahrscheinlichkeitsverteilung</a> überführt.
</p>
<div class="mw-heading mw-heading3"><h3 id="Convolution_Operator">Convolution Operator</h3></div>
<p>Der Convolution Operator ist definiert als <a href="Faltung_(Mathematik)" title="Faltung (Mathematik)">Faltung</a> <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle x*w}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>x</mi>
<mo>∗<!-- ∗ --></mo>
<mi>w</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle x*w}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/7d459bc80876d2b213ba4204c32258c42c276f00.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:5.189ex; height:1.676ex;" alt="{\displaystyle x*w}" loading="lazy"></span> auf den <a href="Reelle_Funktion" class="mw-redirect" title="Reelle Funktion">reellen Funktionen</a> <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle x}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>x</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle x}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/87f9e315fd7e2ba406057a97300593c4802b53e4.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.33ex; height:1.676ex;" alt="{\displaystyle x}" loading="lazy"></span> und <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle w}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>w</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle w}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/88b1e0c8e1be5ebe69d18a8010676fa42d7961e6.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.664ex; height:1.676ex;" alt="{\displaystyle w}" loading="lazy"></span>:<sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup>
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle s(t)=(w*x)(t)=\int x(a)w(t-a)\mathrm {d} a}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>s</mi>
<mo stretchy="false">(</mo>
<mi>t</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo stretchy="false">(</mo>
<mi>w</mi>
<mo>∗<!-- ∗ --></mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">(</mo>
<mi>t</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo>∫<!-- ∫ --></mo>
<mi>x</mi>
<mo stretchy="false">(</mo>
<mi>a</mi>
<mo stretchy="false">)</mo>
<mi>w</mi>
<mo stretchy="false">(</mo>
<mi>t</mi>
<mo>−<!-- − --></mo>
<mi>a</mi>
<mo stretchy="false">)</mo>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">d</mi>
</mrow>
<mi>a</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle s(t)=(w*x)(t)=\int x(a)w(t-a)\mathrm {d} a}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/459d2d1d2decf68f52fee324cc199bb43355edb5.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -2.338ex; width:37.438ex; height:5.676ex;" alt="{\displaystyle s(t)=(w*x)(t)=\int x(a)w(t-a)\mathrm {d} a}" loading="lazy"></span></dd></dl>
<p>Die Zeit wird meistens diskret definiert:
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle s(t)=(w*x)(t)=\sum _{a=-\infty }^{\infty }x(a)w(t-a)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>s</mi>
<mo stretchy="false">(</mo>
<mi>t</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo stretchy="false">(</mo>
<mi>w</mi>
<mo>∗<!-- ∗ --></mo>
<mi>x</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">(</mo>
<mi>t</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<munderover>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>a</mi>
<mo>=</mo>
<mo>−<!-- − --></mo>
<mi mathvariant="normal">∞<!-- ∞ --></mi>
</mrow>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="normal">∞<!-- ∞ --></mi>
</mrow>
</munderover>
<mi>x</mi>
<mo stretchy="false">(</mo>
<mi>a</mi>
<mo stretchy="false">)</mo>
<mi>w</mi>
<mo stretchy="false">(</mo>
<mi>t</mi>
<mo>−<!-- − --></mo>
<mi>a</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle s(t)=(w*x)(t)=\sum _{a=-\infty }^{\infty }x(a)w(t-a)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/28dca7d7b14a380bf9abd9ba952a2dcb40cfe425.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.005ex; width:37.792ex; height:6.843ex;" alt="{\displaystyle s(t)=(w*x)(t)=\sum _{a=-\infty }^{\infty }x(a)w(t-a)}" loading="lazy"></span></dd></dl>
<p>In Anwendungen mit zweidimensionalen Arrays als Input und Kern gilt
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle S(i,j)=(I*K)(i,j)=\sum _{m}\sum _{n}I(m,n)\cdot K(i-m,j-n)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>S</mi>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo stretchy="false">(</mo>
<mi>I</mi>
<mo>∗<!-- ∗ --></mo>
<mi>K</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
</mrow>
</munder>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</munder>
<mi>I</mi>
<mo stretchy="false">(</mo>
<mi>m</mi>
<mo>,</mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
<mo>⋅<!-- ⋅ --></mo>
<mi>K</mi>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>−<!-- − --></mo>
<mi>m</mi>
<mo>,</mo>
<mi>j</mi>
<mo>−<!-- − --></mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle S(i,j)=(I*K)(i,j)=\sum _{m}\sum _{n}I(m,n)\cdot K(i-m,j-n)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/b5834076f3cfdf60b57e7005417388a15b767bf8.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.005ex; width:56.544ex; height:5.509ex;" alt="{\displaystyle S(i,j)=(I*K)(i,j)=\sum _{m}\sum _{n}I(m,n)\cdot K(i-m,j-n)}" loading="lazy"></span></dd></dl>
<p>Der Convolution Operator ist <a href="Kommutativgesetz" title="Kommutativgesetz">kommutativ</a>, d. h. es gilt
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle S(i,j)=(K*I)(i,j)=\sum _{m}\sum _{n}I(i-m,j-n)\cdot K(m,n)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>S</mi>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo stretchy="false">(</mo>
<mi>K</mi>
<mo>∗<!-- ∗ --></mo>
<mi>I</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
</mrow>
</munder>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</munder>
<mi>I</mi>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>−<!-- − --></mo>
<mi>m</mi>
<mo>,</mo>
<mi>j</mi>
<mo>−<!-- − --></mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
<mo>⋅<!-- ⋅ --></mo>
<mi>K</mi>
<mo stretchy="false">(</mo>
<mi>m</mi>
<mo>,</mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle S(i,j)=(K*I)(i,j)=\sum _{m}\sum _{n}I(i-m,j-n)\cdot K(m,n)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/b7e1a2c4e4f225a3f5805dbe33cc9e03109fe36a.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.005ex; width:56.544ex; height:5.509ex;" alt="{\displaystyle S(i,j)=(K*I)(i,j)=\sum _{m}\sum _{n}I(i-m,j-n)\cdot K(m,n)}" loading="lazy"></span></dd></dl>
<p>Außerdem gilt die Cross-Relation:<sup id="cite_ref-17" class="reference"><a href="#cite_note-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup>
</p>
<dl><dd><span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle S(i,j)=(K*I)(i,j)=\sum _{m}\sum _{n}I(i+m,j+n)\cdot K(m,n)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>S</mi>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<mo stretchy="false">(</mo>
<mi>K</mi>
<mo>∗<!-- ∗ --></mo>
<mi>I</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>,</mo>
<mi>j</mi>
<mo stretchy="false">)</mo>
<mo>=</mo>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
</mrow>
</munder>
<munder>
<mo>∑<!-- ∑ --></mo>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</munder>
<mi>I</mi>
<mo stretchy="false">(</mo>
<mi>i</mi>
<mo>+</mo>
<mi>m</mi>
<mo>,</mo>
<mi>j</mi>
<mo>+</mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
<mo>⋅<!-- ⋅ --></mo>
<mi>K</mi>
<mo stretchy="false">(</mo>
<mi>m</mi>
<mo>,</mo>
<mi>n</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle S(i,j)=(K*I)(i,j)=\sum _{m}\sum _{n}I(i+m,j+n)\cdot K(m,n)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/6d8e705067ca2ba900f628e24a6ed3728cda344e.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -3.005ex; width:56.544ex; height:5.509ex;" alt="{\displaystyle S(i,j)=(K*I)(i,j)=\sum _{m}\sum _{n}I(i+m,j+n)\cdot K(m,n)}" loading="lazy"></span></dd></dl>
<div class="mw-heading mw-heading2"><h2 id="Training">Training</h2></div>
<p>CNNs werden in aller Regel <a href="%C3%9Cberwachtes_Lernen" title="Überwachtes Lernen">überwacht</a> trainiert. Während des Trainings wird dabei für jeden gezeigten Input der passende One-Hot-Vektor bereitgestellt. Via <a href="Backpropagation" title="Backpropagation">Backpropagation</a> wird der Gradient eines jeden Neurons berechnet und die Gewichte werden in Richtung des steilsten Abfalls der Fehleroberfläche angepasst.
</p><p>Interessanterweise haben drei vereinfachende Annahmen, die den Berechnungsaufwand des Netzes maßgeblich verringern und damit tiefere Netzwerke zulassen, wesentlich zum Erfolg von CNNs beigetragen.
</p>
<ul><li>Pooling – Hierbei wird der Großteil der Aktivität eines Layers schlicht verworfen.</li>
<li><a href="Rectifier_(neuronale_Netzwerke)" title="Rectifier (neuronale Netzwerke)">ReLU</a> – Die gängige Aktivierungsfunktion, die jeglichen negativen Input auf 0 projiziert.</li>
<li><a href="Dropout_(k%C3%BCnstliches_neuronales_Netz)" title="Dropout (künstliches neuronales Netz)">Dropout</a> – Eine Regularisierungsmethode beim Training, die <a href="%C3%9Cberanpassung" title="Überanpassung">Overfitting</a> verhindert. Dabei werden pro Trainingsschritt zufällig ausgewählte Neuronen aus dem Netzwerk entfernt.</li></ul>
<div class="mw-heading mw-heading2"><h2 id="Expressivität_und_Notwendigkeit"><span id="Expressivit.C3.A4t_und_Notwendigkeit"></span>Expressivität und Notwendigkeit</h2></div>
<p>Da CNNs eine Sonderform von mehrlagigen <a href="Perzeptron" title="Perzeptron">Perzeptrons</a> darstellen,<sup id="cite_ref-LeCun_18-0" class="reference"><a href="#cite_note-LeCun-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup> sind sie prinzipiell identisch in ihrer Ausdrucksstärke.
</p><p>Der Erfolg von CNNs lässt sich mit ihrer kompakten Repräsentation der zu lernenden Gewichte („shared weights“) erklären. Grundlage ist die Annahme, dass ein potentiell interessantes Feature (In Objekterkennung etwa Kanten) an jeder Stelle des Inputsignals (des Bildes) interessant ist. Während ein klassisches zweilagiges Perzeptron mit jeweils 1000 Neuronen pro Ebene für die Verarbeitung von einem Bild im Format 32 × 32 insgesamt mehr als 2 Millionen Gewichte benötigt, verlangt ein CNN mit zwei sich wiederholenden Einheiten, bestehend aus insgesamt 13.000 Neuronen, nur 160.000 (geteilte) zu lernende Gewichte, wovon der Großteil im hinteren Bereich (fully-connected Layer) liegt.
</p><p>Neben dem wesentlich verringerten Arbeitsspeicherbedarf, haben sich geteilte Gewichte als robust gegenüber <a href="Translationsinvarianz" class="mw-redirect" title="Translationsinvarianz">Translations-</a>, Rotations-, <a href="Skaleninvarianz" title="Skaleninvarianz">Skalen-</a> und Luminanzvarianz erwiesen.<sup id="cite_ref-LeCun_18-1" class="reference"><a href="#cite_note-LeCun-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup>
</p><p>Um mithilfe eines mehrlagigen Perzeptrons eine ähnliche Performance in der Bilderkennung zu erreichen, müsste dieses Netzwerk jedes Feature für jeden Bereich des Inputsignals unabhängig erlernen. Dies funktioniert zwar ausreichend für stark verkleinerte Bilder (etwa 32 × 32), aufgrund des <a href="Fluch_der_Dimensionalit%C3%A4t" title="Fluch der Dimensionalität">Fluchs der Dimensionalität</a> scheitern MLPs jedoch an höher auflösenden Bildern.
</p>
<div class="mw-heading mw-heading2"><h2 id="Biologische_Plausibilität"><span id="Biologische_Plausibilit.C3.A4t"></span>Biologische Plausibilität</h2></div>
<p>CNNs können als ein vom <a href="Visueller_Cortex" title="Visueller Cortex">visuellen Cortex</a> inspiriertes Konzept verstanden werden, sind jedoch weit davon entfernt, neuronale Verarbeitung plausibel zu modellieren.
</p><p>Einerseits gilt das Herzstück von CNNs, der Lernmechanismus <a href="Backpropagation" title="Backpropagation">Backpropagation</a>, als biologisch unplausibel, da es bis heute trotz intensiver Bemühungen nicht gelungen ist, neuronale Korrelate von backpropagation-ähnlichen Fehlersignalen zu finden.<sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Biologically_plausible_DL_20-0" class="reference"><a href="#cite_note-Biologically_plausible_DL-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup>
Neben dem stärksten Gegenargument zur biologischen Plausibilität – der Frage, wie der Kortex Zugriff auf das Zielsignal (Label) bekommt – listen Bengio et al. weitere Gründe, darunter die binäre, zeitkontinuierliche Kommunikation biologischer Neurone sowie die Berechnung nicht-linearer Ableitungen der Vorwärtsneuronen<sup id="cite_ref-Biologically_plausible_DL_20-1" class="reference"><a href="#cite_note-Biologically_plausible_DL-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup>.
</p><p>Andererseits konnte durch Untersuchungen mit <a href="Funktionelle_Magnetresonanztomographie" title="Funktionelle Magnetresonanztomographie">fMRT</a> gezeigt werden, dass Aktivierungsmuster einzelner Schichten eines CNNs mit den Neuronenaktivitäten in bestimmten Arealen des visuellen Cortex korrelieren, wenn sowohl das CNN als auch die menschlichen Testprobanden mit ähnlichen Aufgaben aus der Bildverarbeitung konfrontiert werden.<sup id="cite_ref-Neural_encoding_with_DL_21-0" class="reference"><a href="#cite_note-Neural_encoding_with_DL-21"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Modeling_VFWA_with_a_CNN_22-0" class="reference"><a href="#cite_note-Modeling_VFWA_with_a_CNN-22"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup> Neuronen im primären visuellen Cortex, die sogenannten „simple cells“, reagieren auf Aktivität in einem kleinen Bereich der <a href="Netzhaut" title="Netzhaut">Retina</a>. Dieses Verhalten wird in CNNs durch die diskrete Faltung in den convolutional Layers modelliert. Funktional sind diese biologischen Neuronen für die Erkennung von Kanten in bestimmten Orientierungen zuständig. Diese Eigenschaft der simple cells kann wiederum mithilfe von <a href="Gabor-Filter" title="Gabor-Filter">Gabor-Filtern</a> präzise modelliert werden.<sup id="cite_ref-23" class="reference"><a href="#cite_note-23"><span class="cite-bracket">[</span>23<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-24" class="reference"><a href="#cite_note-24"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup> Trainiert man ein CNN zur Objekterkennung, konvergieren die Gewichte im ersten Convolutional Layer ohne jedes „Wissen“ über die Existenz von simple cells gegen Filtermatrizen, die Gabor-Filtern erstaunlich nahe kommen<sup id="cite_ref-25" class="reference"><a href="#cite_note-25"><span class="cite-bracket">[</span>25<span class="cite-bracket">]</span></a></sup>, was als Argument für die biologische Plausibilität von CNNs verstanden werden kann. Angesichts einer umfassenden statistischen Informationsanalyse von Bildern mit dem Ergebnis, dass Ecken und Kanten in verschiedenen Orientierungen die am stärksten voneinander unabhängigen Komponenten in Bildern – und somit die fundamentalsten Grundbausteine zur Bildanalyse – sind, ist dies jedoch zu erwarten.<sup id="cite_ref-26" class="reference"><a href="#cite_note-26"><span class="cite-bracket">[</span>26<span class="cite-bracket">]</span></a></sup>
</p><p>Somit treten die Analogien zwischen Neuronen in CNNs und biologischen Neuronen primär behavioristisch zutage, also im Vergleich zweier funktionsfähiger Systeme, wohingegen die Entwicklung eines „unwissenden“ Neurons zu einem (beispielsweise) gesichtserkennenden Neuron in beiden Systemen diametralen Prinzipien folgt.
</p>
<div class="mw-heading mw-heading2"><h2 id="Probleme">Probleme</h2></div>
<p>Neuronale Feedforward Netze können jede <a href="Stetige_Funktion" title="Stetige Funktion">stetige Funktion</a> der Form <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle f:X\in \mathbb {R} ^{n}\to Y\in \mathbb {R} ^{m}}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>f</mi>
<mo>:</mo>
<mi>X</mi>
<mo>∈<!-- ∈ --></mo>
<msup>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="double-struck">R</mi>
</mrow>
<mrow class="MJX-TeXAtom-ORD">
<mi>n</mi>
</mrow>
</msup>
<mo stretchy="false">→<!-- → --></mo>
<mi>Y</mi>
<mo>∈<!-- ∈ --></mo>
<msup>
<mrow class="MJX-TeXAtom-ORD">
<mi mathvariant="double-struck">R</mi>
</mrow>
<mrow class="MJX-TeXAtom-ORD">
<mi>m</mi>
</mrow>
</msup>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle f:X\in \mathbb {R} ^{n}\to Y\in \mathbb {R} ^{m}}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/b7242804851ee0fbb59b2a52b26e275420bde894.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.671ex; width:22.514ex; height:2.676ex;" alt="{\displaystyle f:X\in \mathbb {R} ^{n}\to Y\in \mathbb {R} ^{m}}" loading="lazy"></span> annähern, dies wird durch das universelle Approximationstheorem garantiert. Es gibt jedoch keine Garantie dafür, dass das Training dies auch ermöglicht. Wird während dem Training keine Regularisierung verwendet, wird ein neuronales Netz <a href="%C3%9Cberanpassung" title="Überanpassung">überangepasst</a> auf das Rauschen in den Trainingsdaten. Durch die Verringerung der Modellkapazität durch Wiederverwendung der Gewichte mithilfe einer <a href="Faltung_(Mathematik)" title="Faltung (Mathematik)">Faltung</a> (als eine bestimmte Art der Regularisierung) kann die Tendenz zur Überanpassung verringert werden.
Formal bedeutet dies, dass die Merkmale <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X\to X'}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
<mo stretchy="false">→<!-- → --></mo>
<msup>
<mi>X</mi>
<mo>′</mo>
</msup>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X\to X'}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/ce3b51190a5a80ac3ff36f45611676482388b60a.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:8.276ex; height:2.509ex;" alt="{\displaystyle X\to X'}" loading="lazy"></span> transformiert werden, so dass, wenn <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle H(X)}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>H</mi>
<mo stretchy="false">(</mo>
<mi>X</mi>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle H(X)}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/bd232b6fb5ea803efc1154d2efb0c3fe00a4531b.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:5.853ex; height:2.843ex;" alt="{\displaystyle H(X)}" loading="lazy"></span> die <a href="Entropie_(Informationstheorie)" title="Entropie (Informationstheorie)">Entropie</a> von <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle X}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>X</mi>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle X}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/68baa052181f707c662844a465bfeeb135e82bab.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.338ex; width:1.98ex; height:2.176ex;" alt="{\displaystyle X}" loading="lazy"></span> ist, dann gilt <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle H(X)\leq H(X')}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>H</mi>
<mo stretchy="false">(</mo>
<mi>X</mi>
<mo stretchy="false">)</mo>
<mo>≤<!-- ≤ --></mo>
<mi>H</mi>
<mo stretchy="false">(</mo>
<msup>
<mi>X</mi>
<mo>′</mo>
</msup>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle H(X)\leq H(X')}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/810065ce7755a3f6c4a62c8e59ec1c839ac039ac.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:15.506ex; height:3.009ex;" alt="{\displaystyle H(X)\leq H(X')}" loading="lazy"></span>. Dadurch kann das neuronale Netz eine neue Zielfunktion <span class="mwe-math-element mwe-math-element-inline"><span class="mwe-math-mathml-inline mwe-math-mathml-a11y" style="display: none;"><math xmlns="http://www.w3.org/1998/Math/MathML" alttext="{\displaystyle g:f(X)\to g(X')}">
<semantics>
<mrow class="MJX-TeXAtom-ORD">
<mstyle displaystyle="true" scriptlevel="0">
<mi>g</mi>
<mo>:</mo>
<mi>f</mi>
<mo stretchy="false">(</mo>
<mi>X</mi>
<mo stretchy="false">)</mo>
<mo stretchy="false">→<!-- → --></mo>
<mi>g</mi>
<mo stretchy="false">(</mo>
<msup>
<mi>X</mi>
<mo>′</mo>
</msup>
<mo stretchy="false">)</mo>
</mstyle>
</mrow>
<annotation encoding="application/x-tex">{\displaystyle g:f(X)\to g(X')}</annotation>
</semantics>
</math></span><img src="./_assets_/eb734a37dd21ce173a46342d1cc64c92/2a1903490ed526b87a1c0144db3dc1c6e07e4500.svg" class="mwe-math-fallback-image-inline mw-invert skin-invert" aria-hidden="true" style="vertical-align: -0.838ex; width:17.342ex; height:3.009ex;" alt="{\displaystyle g:f(X)\to g(X')}" loading="lazy"></span> lernen, die weniger anfällig für Überanpassung ist. Dies wird dadurch ermöglicht, dass eine wichtige Voraussetzung, die <a href="Lineare_Unabh%C3%A4ngigkeit" title="Lineare Unabhängigkeit">lineare Unabhängigkeit</a> der Merkmale, für einige Datenklassen systematisch verletzt wird.<sup id="cite_ref-27" class="reference"><a href="#cite_note-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Anwendung">Anwendung</h2></div>
<p>Seit dem Einsatz von <a href="Grafikprozessor" title="Grafikprozessor">Grafikprozessor</a>-Programmierung können CNNs erstmals effizient trainiert werden.<sup id="cite_ref-28" class="reference"><a href="#cite_note-28"><span class="cite-bracket">[</span>28<span class="cite-bracket">]</span></a></sup> Sie gelten als <a href="State_of_the_Art" class="mw-redirect" title="State of the Art">State-of-the-Art</a>-Methode für zahlreiche Anwendungen im Bereich der Klassifizierung.
</p>
<div class="mw-heading mw-heading3"><h3 id="Bilderkennung">Bilderkennung</h3></div>
<p>Im Jahr 2012 verbesserte das CNN <a href="AlexNet" title="AlexNet">AlexNet</a> die Fehlerquote beim jährlichen Wettbewerb der Benchmark-Datenbank <a href="ImageNet" title="ImageNet">ImageNet</a> (ILSVRC) von dem vormaligen Rekord von 25,8&nbsp;% auf 16,4&nbsp;%. Seitdem nutzen alle vorne platzierten Algorithmen CNN-Strukturen.
</p><p>Im Jahr 2016 wurde schon eine Fehlerquote &lt; 3&nbsp;% erreicht.<sup id="cite_ref-29" class="reference"><a href="#cite_note-29"><span class="cite-bracket">[</span>29<span class="cite-bracket">]</span></a></sup> Auf der am häufigsten genutzten Bilddatenbanken <a href="MNIST-Datenbank" title="MNIST-Datenbank">MNIST</a> wurde im selben Jahr eine Fehlerquote von 0,23&nbsp;% erreicht, was der geringsten Fehlerquote aller jemals getesteten Algorithmen entsprach.<sup id="cite_ref-mcdns_30-0" class="reference"><a href="#cite_note-mcdns-30"><span class="cite-bracket">[</span>30<span class="cite-bracket">]</span></a></sup>
</p><p>Auch im Bereich der Gesichtserkennung konnten bahnbrechende Resultate erzielt werden.<sup id="cite_ref-31" class="reference"><a href="#cite_note-31"><span class="cite-bracket">[</span>31<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Spracherkennung">Spracherkennung</h3></div>
<p>CNNs werden erfolgreich zur <a href="Spracherkennung" title="Spracherkennung">Spracherkennung</a> eingesetzt und haben hervorragende Resultate in folgenden Bereichen erzielt:
</p>
<ul><li>semantisches Parsen<sup id="cite_ref-32" class="reference"><a href="#cite_note-32"><span class="cite-bracket">[</span>32<span class="cite-bracket">]</span></a></sup></li>
<li>Suchanfragenrückerkennung<sup id="cite_ref-33" class="reference"><a href="#cite_note-33"><span class="cite-bracket">[</span>33<span class="cite-bracket">]</span></a></sup></li>
<li>Satzmodellierung<sup id="cite_ref-34" class="reference"><a href="#cite_note-34"><span class="cite-bracket">[</span>34<span class="cite-bracket">]</span></a></sup></li>
<li>Satzklassifizierung<sup id="cite_ref-35" class="reference"><a href="#cite_note-35"><span class="cite-bracket">[</span>35<span class="cite-bracket">]</span></a></sup></li>
<li><a href="Part-of-speech-Tagging" title="Part-of-speech-Tagging">Part-of-speech-Tagging</a><sup id="cite_ref-36" class="reference"><a href="#cite_note-36"><span class="cite-bracket">[</span>36<span class="cite-bracket">]</span></a></sup></li>
<li><a href="Maschinelle_%C3%9Cbersetzung" title="Maschinelle Übersetzung">Maschinelle Übersetzung</a> (z.&nbsp;B. verwendet im Online-Dienst <a href="DeepL" title="DeepL">DeepL</a>)<sup id="cite_ref-37" class="reference"><a href="#cite_note-37"><span class="cite-bracket">[</span>37<span class="cite-bracket">]</span></a></sup></li></ul>
<div class="mw-heading mw-heading3"><h3 id="Reinforcement_Learning">Reinforcement Learning</h3></div>
<p>Angewendet werden können CNNs auch im Bereich <a href="Best%C3%A4rkendes_Lernen" title="Bestärkendes Lernen">Reinforcement Learning</a>, bei dem ein CNN mit <a href="Q-Learning" class="mw-redirect" title="Q-Learning">Q-Learning</a> kombiniert wird. Das Netzwerk wird darauf trainiert zu schätzen, welche Aktionen bei einem gegebenen Zustand zu welchem zukünftigen Gewinn führen. Durch die Verwendung eines CNNs können so auch komplexe, höher-dimensionale Zustandsräume betrachtet werden, wie etwa die Bildschirmausgabe eines <a href="Videospiel" class="mw-redirect" title="Videospiel">Videospiels</a>.<sup id="cite_ref-38" class="reference"><a href="#cite_note-38"><span class="cite-bracket">[</span>38<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Landwirtschaft">Landwirtschaft</h3></div>
<p>CNNs werden heutzutage auch eingesetzt, um die Auswirkungen von Düngung vorherzusagen und Düngungsempfehlungen abzugeben.<sup id="cite_ref-rani2022_39-0" class="reference"><a href="#cite_note-rani2022-39"><span class="cite-bracket">[</span>39<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-sethy2021_40-0" class="reference"><a href="#cite_note-sethy2021-40"><span class="cite-bracket">[</span>40<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Literatur">Literatur</h2></div>
<ul><li><a href="Ian_Goodfellow" title="Ian Goodfellow">Ian Goodfellow</a>, Yoshua Bengio, Aaron Courville: <cite style="font-style:italic">Deep Learning</cite> (=&nbsp;<cite style="font-style:italic">Adaptive Computation and Machine Learning</cite>). MIT Press, 2016, ISBN 978-0-262-03561-3, 9 Convolutional Networks (<a rel="nofollow" class="external text" href="http://www.deeplearningbook.org/">Online</a>).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abookitem&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=9+Convolutional+Networks&amp;rft.au=Ian+Goodfellow%2C+Yoshua+Bengio%2C+Aaron+Courville&amp;rft.btitle=Deep+Learning&amp;rft.date=2016&amp;rft.genre=bookitem&amp;rft.isbn=9780262035613&amp;rft.pub=MIT+Press&amp;rft.series=Adaptive+Computation+and+Machine+Learning" style="display:none">&nbsp;</span></li></ul>
<div class="mw-heading mw-heading2"><h2 id="Weblinks">Weblinks</h2></div>
<ul><li><a rel="nofollow" class="external text" href="https://www.ted.com/talks/fei_fei_li_how_we_re_teaching_computers_to_understand_pictures">TED-Talk: <i>How we are teaching computers to understand pictures – Fei Fei Li,</i></a> März 2015, abgerufen am 17. November 2016.</li>
<li><a rel="nofollow" class="external text" href="http://scs.ryerson.ca/~aharley/vis/conv/flat.html">2D-Visualisierung der Aktivität eines zweilagigen CNNs,</a> abgerufen am 17. November 2016.</li>
<li><a rel="nofollow" class="external text" href="https://github.com/Hvass-Labs/TensorFlow-Tutorials/blob/master/02_Convolutional_Neural_Network.ipynb">Tutorial zur Implementierung eines CNN</a> mithilfe der <a href="Python_(Programmiersprache)" title="Python (Programmiersprache)">Python</a>-Bibliothek <a href="TensorFlow" title="TensorFlow">TensorFlow</a></li>
<li><a rel="nofollow" class="external text" href="http://cs231n.github.io/convolutional-networks/">CNN-Tutorial der University of Stanford,</a> inklusive Visualisierung erlernter <a href="Faltungsmatrix" title="Faltungsmatrix">Faltungsmatrizen</a>, abgerufen am 17. November 2016.</li>
<li><a rel="nofollow" class="external text" href="http://yann.lecun.com/exdb/publis/pdf/lecun-98.pdf"><i>Gradient-Based Learning Applied to Document Recognition, Y. Le Cun et al</i></a> (PDF; 933&nbsp;kB), erste erfolgreiche Anwendung eines CNN, abgerufen am 17. November 2016.</li>
<li><a rel="nofollow" class="external text" href="https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf"><i>ImageNet Classification with Deep Convolutional Neural Networks, A. Krizhevsky, I. Sutskever and G. E. Hinton,</i></a> AlexNet – Durchbruch in der Bilderkennung, Gewinner der ILSVRC-Challenge 2012.</li></ul>
<div class="mw-heading mw-heading2"><h2 id="Einzelnachweise">Einzelnachweise</h2></div>
<ol class="references">
<li id="cite_note-robust_face_detection-1"><span class="mw-cite-backlink"><a href="#cite_ref-robust_face_detection_1-0">↑</a></span> <span class="reference-text">Masakazu Matsugu, Katsuhiko Mori, Yusuke Mitari, Yuji Kaneda: <cite style="font-style:italic">Subject independent facial expression recognition with robust face detection using a convolutional neural network</cite>. In: <cite style="font-style:italic">Neural Networks</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em">&nbsp;</span>16</span>, <span style="white-space:nowrap">Nr.<span style="display:inline-block;width:.2em">&nbsp;</span>5</span>, 2003, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>555–559</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1016/S0893-6080%2803%2900115-1">10.1016/S0893-6080(03)00115-1</a></span> (<a rel="nofollow" class="external text" href="http://www.iro.umontreal.ca/~pift6080/H09/documents/papers/sparse/matsugo_etal_face_expression_conv_nnet.pdf">Online</a> [PDF; abgerufen am 28.&nbsp;Mai 2017]).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Ajournal&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Subject+independent+facial+expression+recognition+with+robust+face+detection+using+a+convolutional+neural+network&amp;rft.au=Masakazu+Matsugu%2C+Katsuhiko+Mori%2C+Yusuke+Mitari%2C+...&amp;rft.date=2003&amp;rft.doi=10.1016%2FS0893-6080%2803%2900115-1&amp;rft.genre=journal&amp;rft.issue=5&amp;rft.jtitle=Neural+Networks&amp;rft.pages=555-559&amp;rft.volume=16" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-fukushima1980-2"><span class="mw-cite-backlink"><a href="#cite_ref-fukushima1980_2-0">↑</a></span> <span class="reference-text">Kunihiko Fukushima: <cite style="font-style:italic">Neocognitron: A Self-organizing Neural Network Model for a Mechanism of Pattern Recognition Unaffected by Shift in Position</cite>. In: <cite style="font-style:italic">Biological Cybernetics</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em">&nbsp;</span>36</span>, <span style="white-space:nowrap">Nr.<span style="display:inline-block;width:.2em">&nbsp;</span>4</span>, 1980, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>193–202</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1007/BF00344251">10.1007/BF00344251</a></span> (<a rel="nofollow" class="external text" href="https://www.cs.princeton.edu/courses/archive/spr08/cos598B/Readings/Fukushima1980.pdf">Online</a> [PDF]).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Ajournal&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Neocognitron%3A+A+Self-organizing+Neural+Network+Model+for+a+Mechanism+of+Pattern+Recognition+Unaffected+by+Shift+in+Position&amp;rft.au=Kunihiko+Fukushima&amp;rft.date=1980&amp;rft.doi=10.1007%2FBF00344251&amp;rft.genre=journal&amp;rft.issue=4&amp;rft.jtitle=Biological+Cybernetics&amp;rft.pages=193-202&amp;rft.volume=36" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-schmidhuber2015-3"><span class="mw-cite-backlink"><a href="#cite_ref-schmidhuber2015_3-0">↑</a></span> <span class="reference-text">Jürgen Schmidhuber: <cite style="font-style:italic">Deep Learning in Neural Networks: An Overview</cite>. In: <cite style="font-style:italic">Neural Networks</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em">&nbsp;</span>61</span>, 2015, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>85–117</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1016/j.neunet.2014.09.003">10.1016/j.neunet.2014.09.003</a></span>, <a href="ArXiv" title="ArXiv">arxiv</a>:<a rel="nofollow" class="external text" href="https://arxiv.org/abs/1404.7828">1404.7828</a>.<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Deep+Learning+in+Neural+Networks%3A+An+Overview&amp;rft.au=J%C3%BCrgen+Schmidhuber&amp;rft.btitle=Neural+Networks&amp;rft.date=2015&amp;rft.doi=10.1016%2Fj.neunet.2014.09.003&amp;rft.genre=book&amp;rft.pages=85-117&amp;rft.volume=61" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-Waibel1987-4"><span class="mw-cite-backlink"><a href="#cite_ref-Waibel1987_4-0">↑</a></span> <span class="reference-text">Alex Waibel: <cite class="lang" lang="en" dir="auto" style="font-style:italic">Phoneme Recognition Using Time-Delay Neural Networks</cite>. Meeting of the Institute of Electrical, Information and Communication Engineers (IEICE), Tokyo, Japan, 1987. 1987 (englisch).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.au=Alex%26%2332%3BWaibel&amp;rft.btitle=Phoneme+Recognition+Using+Time-Delay+Neural+Networks&amp;rft.date=1987&amp;rft.genre=book" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-lecun1989-5"><span class="mw-cite-backlink"><a href="#cite_ref-lecun1989_5-0">↑</a></span> <span class="reference-text">Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, L. D. Jackel, <a rel="nofollow" class="external text" href="http://yann.lecun.com/exdb/publis/pdf/lecun-89e.pdf">Backpropagation Applied to Handwritten Zip Code Recognition</a>; AT&amp;T Bell Laboratories, 1989</span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><a href="#cite_ref-6">↑</a></span> <span class="reference-text">Yann LeCun, Leon Bottou, Yoshua Bengio, Patrick Haffner: <cite style="font-style:italic">Gradient-based Learning Applied to Document Recognition</cite>. In: <cite style="font-style:italic">Proceedings of the IEEE</cite>. 1998 (<a rel="nofollow" class="external text" href="http://yann.lecun.com/exdb/publis/pdf/lecun-01a.pdf">lecun.com</a> [PDF]).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Gradient-based+Learning+Applied+to+Document+Recognition&amp;rft.au=Yann+LeCun%2C+Leon+Bottou%2C+Yoshua+Bengio%2C+...&amp;rft.btitle=Proceedings+of+the+IEEE&amp;rft.date=1998&amp;rft.genre=book" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><a href="#cite_ref-7">↑</a></span> <span class="reference-text">Convolutional Neural Networks Demystified: A Matched Filtering Perspective Based Tutorial <a rel="nofollow" class="external free" href="https://arxiv.org/abs/2108.11663v3">https://arxiv.org/abs/2108.11663v3</a></span>
</li>
<li id="cite_note-deeplearning-8"><span class="mw-cite-backlink"><a href="#cite_ref-deeplearning_8-0">↑</a></span> <span class="reference-text"><span class="cite">unknown: <a rel="nofollow" class="external text" href="https://web.archive.org/web/20171228091645/http://deeplearning.net/tutorial/lenet.html"><i>Convolutional Neural Networks (LeNet).</i></a> Archiviert vom <style data-mw-deduplicate="TemplateStyles:r250917974">
/* start https://de.wikipedia.org/ */


.mw-parser-output .dewiki-iconexternal>a{background-position:center right!important;background-repeat:no-repeat!important}body.skin-minerva .mw-parser-output .dewiki-iconexternal>a{background-image:url("./_mw_/OOjs_UI_icon_external-link-ltr-progressive.svg")!important;background-size:10px!important;padding-right:13px!important}body.skin-timeless .mw-parser-output .dewiki-iconexternal>a,body.skin-monobook .mw-parser-output .dewiki-iconexternal>a{background-image:url("./_mw_/MediaWiki_external_link_icon.svg")!important;padding-right:13px!important}body.skin-vector .mw-parser-output .dewiki-iconexternal>a{background-image:url("./_mw_/Link.ernal-small-ltr-progressive.svg")!important;background-size:0.857em!important;padding-right:1em!important}


/* end https://de.wikipedia.org/ */
</style><span class="dewiki-iconexternal"><a class="external text" href="https://redirecter.toolforge.org/?url=http%3A%2F%2Fdeeplearning.net%2Ftutorial%2Flenet.html">Original</a></span> (nicht mehr online verfügbar) am <span style="white-space:nowrap;">28.&nbsp;Dezember 2017</span><span>;</span><span class="Abrufdatum"> abgerufen am 17.&nbsp;November 2016</span> (englisch).</span> <small class="archiv-bot"><span class="wp_boppel noviewer" aria-hidden="true" role="presentation"><span typeof="mw:File"><span title="i"></span></span></span>&nbsp;<b>Info:</b> Der Archivlink wurde automatisch eingesetzt und noch nicht geprüft. Bitte prüfe Original- und Archivlink gemäß Anleitung und entferne dann diesen Hinweis.</small><span style="display:none"><a rel="nofollow" class="external text" href="http://IABotmemento.invalid/http://deeplearning.net/tutorial/lenet.html">@1</a></span><span style="display:none"><a rel="nofollow" class="external text" href="http://deeplearning.net/tutorial/lenet.html">@2</a></span><span style="display:none">Vorlage:Webachiv/IABot/deeplearning.net</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=Convolutional+Neural+Networks+%28LeNet%29&amp;rft.description=Convolutional+Neural+Networks+%28LeNet%29&amp;rft.identifier=https%3A%2F%2Fweb.archive.org%2Fweb%2F20171228091645%2Fhttp%3A%2F%2Fdeeplearning.net%2Ftutorial%2Flenet.html&amp;rft.creator=unknown&amp;rft.source=http://deeplearning.net/tutorial/lenet.html&amp;rft.language=en">&nbsp;</span></span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><a href="#cite_ref-9">↑</a></span> <span class="reference-text">D. H. Hubel, T. N. Wiesel: <cite style="font-style:italic">Receptive fields and functional architecture of monkey striate cortex</cite>. In: <cite style="font-style:italic">The Journal of Physiology</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em">&nbsp;</span>195</span>, <span style="white-space:nowrap">Nr.<span style="display:inline-block;width:.2em">&nbsp;</span>1</span>, 1.&nbsp;März 1968, <a href="Internationale_Standardnummer_f%C3%BCr_fortlaufende_Sammelwerke" title="Internationale Standardnummer für fortlaufende Sammelwerke">ISSN</a>&nbsp;<span style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://zdb-katalog.de/list.xhtml?t=iss%3D%220022-3751%22&amp;key=cql">0022-3751</a></span>, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>215–243</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1113/jphysiol.1968.sp008455">10.1113/jphysiol.1968.sp008455</a></span>, <a class="external mw-magiclink-pmid" rel="nofollow" href="https://www.ncbi.nlm.nih.gov/pubmed/4966457?dopt=Abstract">PMID 4966457</a>, <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1557912/">PMC&nbsp;1557912</a> (freier Volltext).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Ajournal&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Receptive+fields+and+functional+architecture+of+monkey+striate+cortex&amp;rft.au=D.+H.+Hubel%2C+T.+N.+Wiesel&amp;rft.date=1968-03-01&amp;rft.doi=10.1113%2Fjphysiol.1968.sp008455&amp;rft.genre=journal&amp;rft.issn=0022-3751&amp;rft.issue=1&amp;rft.jtitle=The+Journal+of+Physiology&amp;rft.pages=215-243&amp;rft.pmc=1557912&amp;rft.pmid=4966457&amp;rft.volume=195" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-10"><span class="mw-cite-backlink"><a href="#cite_ref-10">↑</a></span> <span class="reference-text">Jason Brownlee, Machine Learning Mastery: <a rel="nofollow" class="external text" href="https://machinelearningmastery.com/convolutional-layers-for-deep-learning-neural-networks/"><i>How Do Convolutional Layers Work in Deep Learning Neural Networks?</i></a></span>
</li>
<li id="cite_note-:0-11"><span class="mw-cite-backlink">↑ <sup><a href="#cite_ref-:0_11-0">a</a></sup> <sup><a href="#cite_ref-:0_11-1">b</a></sup></span> <span class="reference-text">Mayank Mishra, Towards Data Science: <a rel="nofollow" class="external text" href="https://towardsdatascience.com/convolutional-neural-networks-explained-9cc5188c4939"><i>Convolutional Neural Networks, Explained</i></a></span>
</li>
<li id="cite_note-weng1993-12"><span class="mw-cite-backlink"><a href="#cite_ref-weng1993_12-0">↑</a></span> <span class="reference-text">J Weng, N Ahuja, TS Huang: <cite style="font-style:italic">Learning recognition and segmentation of 3-D objects from 2-D images</cite>. In: <cite style="font-style:italic">Proc. 4th International Conf. Computer Vision</cite>. 1993, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>121–128</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1109/ICCV.1993.378228">10.1109/ICCV.1993.378228</a></span>.<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Learning+recognition+and+segmentation+of+3-D+objects+from+2-D+images&amp;rft.au=J+Weng%2C+N+Ahuja%2C+TS+Huang&amp;rft.btitle=Proc.+4th+International+Conf.+Computer+Vision&amp;rft.date=1993&amp;rft.doi=10.1109%2FICCV.1993.378228&amp;rft.genre=book&amp;rft.pages=121-128" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-Yamaguchi1990-13"><span class="mw-cite-backlink"><a href="#cite_ref-Yamaguchi1990_13-0">↑</a></span> <span class="reference-text">Kouichi Yamaguchi, Kenji Sakamoto, Toshio Sakamoto, Yoshiji Fujimoto: <cite class="lang" lang="en" dir="auto" style="font-style:italic">A Neural Network for Speaker-Independent Isolated Word Recognition</cite>. First International Conference on Spoken Language Processing (ICSLP 1990). Kobe, Japan November 1990 (englisch, <a rel="nofollow" class="external text" href="https://web.archive.org/web/20210307233750/https://www.isca-speech.org/archive/icslp_1990/i90_1077.html">isca-speech.org</a> (<span class="webarchiv-memento"><a href="Webarchivierung#Begrifflichkeiten" title="Webarchivierung">Memento</a></span> des <span class="dewiki-iconexternal"><a class="external text" href="https://redirecter.toolforge.org/?url=https%3A%2F%2Fwww.isca-speech.org%2Farchive%2Ficslp_1990%2Fi90_1077.html">Originals</a></span> vom 7.&nbsp;März 2021 im <i><a href="Internet_Archive" title="Internet Archive">Internet Archive</a></i>) [abgerufen am 13.&nbsp;August 2022]).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.au=Kouichi%26%2332%3BYamaguchi%2C%26%2332%3BKenji%26%2332%3BSakamoto%2C%26%2332%3BToshio%26%2332%3BSakamoto%2C+...&amp;rft.btitle=A+Neural+Network+for+Speaker-Independent+Isolated+Word+Recognition&amp;rft.date=1990-11&amp;rft.genre=book&amp;rft.place=Kobe%2C+Japan" style="display:none">&nbsp;</span> <small class="archiv-bot"><span class="wp_boppel noviewer" aria-hidden="true" role="presentation"><span typeof="mw:File"><span title="i"></span></span></span>&nbsp;<b>Info:</b> Der Archivlink wurde automatisch eingesetzt und noch nicht geprüft. Bitte prüfe Original- und Archivlink gemäß Anleitung und entferne dann diesen Hinweis.</small><span style="display:none"><a rel="nofollow" class="external text" href="http://IABotmemento.invalid/https://www.isca-speech.org/archive/icslp_1990/i90_1077.html">@1</a></span><span style="display:none"><a rel="nofollow" class="external text" href="https://www.isca-speech.org/archive/icslp_1990/i90_1077.html">@2</a></span><span style="display:none">Vorlage:Webachiv/IABot/www.isca-speech.org</span></span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><a href="#cite_ref-14">↑</a></span> <span class="reference-text">Benjamin Graham: <cite style="font-style:italic">Fractional Max-Pooling</cite>. 18.&nbsp;Dezember 2014, <a href="ArXiv" title="ArXiv">arxiv</a>:<a rel="nofollow" class="external text" href="https://arxiv.org/abs/1412.6071">1412.6071</a>.<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.au=Benjamin+Graham&amp;rft.btitle=Fractional+Max-Pooling&amp;rft.date=2014-12-18&amp;rft.genre=book" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-Scherer-15"><span class="mw-cite-backlink"><a href="#cite_ref-Scherer_15-0">↑</a></span> <span class="reference-text">Dominik Scherer, Andreas C. Müller, Sven Müller: <cite class="lang" lang="en" dir="auto" style="font-style:italic">Evaluation of Pooling Operations in Convolutional Architectures for Object Recognition</cite>. Artificial Neural Networks (ICANN), 20th International Conference on. Springer, Thessaloniki, Greece 2010, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>92–101</span> (englisch, <a rel="nofollow" class="external text" href="http://ais.uni-bonn.de/papers/icann2010_maxpool.pdf">uni-bonn.de</a> [PDF]).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.au=Dominik%26%2332%3BScherer%2C%26%2332%3BAndreas+C.%26%2332%3BM%C3%BCller%2C%26%2332%3BSven%26%2332%3BM%C3%BCller&amp;rft.btitle=Evaluation+of+Pooling+Operations+in+Convolutional+Architectures+for+Object+Recognition&amp;rft.date=2010&amp;rft.genre=book&amp;rft.pages=92-101&amp;rft.place=Thessaloniki%2C+Greece&amp;rft.pub=Springer" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><a href="#cite_ref-16">↑</a></span> <span class="reference-text">Josif Grabocka, Stiftung Universität Hildesheim: <a rel="nofollow" class="external text" href="https://www.ismll.uni-hildesheim.de/lehre/dl-17s/script/2017-05-31-cnn.pdf"><i>Convolutional Neural Networks (CNN)</i></a></span>
</li>
<li id="cite_note-17"><span class="mw-cite-backlink"><a href="#cite_ref-17">↑</a></span> <span class="reference-text">Zoran Nikolić, Universität zu Köln: <a rel="nofollow" class="external text" href="https://www.mi.uni-koeln.de/wp-znikolic/wp-content/uploads/2019/06/11-Odenthal.pdf"><i>Convolutional Neural Networks</i></a></span>
</li>
<li id="cite_note-LeCun-18"><span class="mw-cite-backlink">↑ <sup><a href="#cite_ref-LeCun_18-0">a</a></sup> <sup><a href="#cite_ref-LeCun_18-1">b</a></sup></span> <span class="reference-text">
<span class="cite">Yann LeCun: <a rel="nofollow" class="external text" href="http://yann.lecun.com/exdb/lenet/"><i>LeNet-5, convolutional neural networks.</i></a><span class="Abrufdatum"> Abgerufen am 17.&nbsp;November 2016</span>.</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=LeNet-5%2C+convolutional+neural+networks&amp;rft.description=LeNet-5%2C+convolutional+neural+networks&amp;rft.identifier=http%3A%2F%2Fyann.lecun.com%2Fexdb%2Flenet%2F&amp;rft.creator=Yann+LeCun">&nbsp;</span></span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><a href="#cite_ref-19">↑</a></span> <span class="reference-text">P. Mazzoni, R. A. Andersen, M. I. Jordan: <i>A more biologically plausible learning rule than backpropagation applied to a network model of cortical area 7a.</i> In: <i>Cerebral cortex.</i> Band 1, Nummer 4, 1991 Jul-Aug, S.&nbsp;293–307, <a href="https://doi.org/10.1093/cercor/1.4.293" class="extiw external" title="doi:10.1093/cercor/1.4.293">doi:10.1093/cercor/1.4.293</a>, <a class="external mw-magiclink-pmid" rel="nofollow" href="https://www.ncbi.nlm.nih.gov/pubmed/1822737?dopt=Abstract">PMID 1822737</a>.</span>
</li>
<li id="cite_note-Biologically_plausible_DL-20"><span class="mw-cite-backlink">↑ <sup><a href="#cite_ref-Biologically_plausible_DL_20-0">a</a></sup> <sup><a href="#cite_ref-Biologically_plausible_DL_20-1">b</a></sup></span> <span class="reference-text">
Yoshua Bengio: <cite class="lang" lang="en" dir="auto" style="font-style:italic">Towards Biologically Plausible Deep Learning</cite>. Februar 2015, <a href="ArXiv" title="ArXiv">arxiv</a>:<a rel="nofollow" class="external text" href="https://arxiv.org/abs/1502.04156v3">1502.04156v3</a> (englisch).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.au=Yoshua+Bengio&amp;rft.btitle=Towards+Biologically+Plausible+Deep+Learning&amp;rft.date=2015-02&amp;rft.genre=book" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-Neural_encoding_with_DL-21"><span class="mw-cite-backlink"><a href="#cite_ref-Neural_encoding_with_DL_21-0">↑</a></span> <span class="reference-text">
Haiguang Wen: <cite class="lang" lang="en" dir="auto" style="font-style:italic">Neural Encoding and Decoding with Deep Learning for Dynamic Natural Vision</cite>. August 2016, <a href="ArXiv" title="ArXiv">arxiv</a>:<a rel="nofollow" class="external text" href="https://arxiv.org/abs/1608.03425">1608.03425</a> (englisch).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.au=Haiguang+Wen&amp;rft.btitle=Neural+Encoding+and+Decoding+with+Deep+Learning+for+Dynamic+Natural+Vision&amp;rft.date=2016-08&amp;rft.genre=book" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-Modeling_VFWA_with_a_CNN-22"><span class="mw-cite-backlink"><a href="#cite_ref-Modeling_VFWA_with_a_CNN_22-0">↑</a></span> <span class="reference-text">
<span class="cite">Sandy Wiraatmadja: <a rel="nofollow" class="external text" href="https://mindmodeling.org/cogsci2016/papers/0435/paper0435.pdf"><i>Modeling the Visual Word Form Area Using a Deep Convolutional Neural Network.</i></a> (PDF)<span class="Abrufdatum"> Abgerufen am 17.&nbsp;September 2017</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=Modeling+the+Visual+Word+Form+Area+Using+a+Deep+Convolutional+Neural+Network&amp;rft.description=Modeling+the+Visual+Word+Form+Area+Using+a+Deep+Convolutional+Neural+Network&amp;rft.identifier=https%3A%2F%2Fmindmodeling.org%2Fcogsci2016%2Fpapers%2F0435%2Fpaper0435.pdf&amp;rft.creator=Sandy+Wiraatmadja&amp;rft.language=en">&nbsp;</span></span>
</li>
<li id="cite_note-23"><span class="mw-cite-backlink"><a href="#cite_ref-23">↑</a></span> <span class="reference-text"><a href="John_Daugman" title="John Daugman">J. G. Daugman</a>: <i>Uncertainty relation for resolution in space, spatial frequency, and orientation optimized by two-dimensional visual cortical filters.</i> In: <i>Journal of the Optical Society of America A,</i> 2 (7): 1160–1169, July 1985.</span>
</li>
<li id="cite_note-24"><span class="mw-cite-backlink"><a href="#cite_ref-24">↑</a></span> <span class="reference-text">S. Marčelja: <cite style="font-style:italic">Mathematical description of the responses of simple cortical cells</cite>. In: <cite style="font-style:italic">Journal of the Optical Society of America</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em">&nbsp;</span>70</span>, <span style="white-space:nowrap">Nr.<span style="display:inline-block;width:.2em">&nbsp;</span>11</span>, 1980, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>1297–1300</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1364/JOSA.70.001297">10.1364/JOSA.70.001297</a></span>.<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Ajournal&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Mathematical+description+of+the+responses+of+simple+cortical+cells&amp;rft.au=S.+Mar%C4%8Delja&amp;rft.date=1980&amp;rft.doi=10.1364%2FJOSA.70.001297&amp;rft.genre=journal&amp;rft.issue=11&amp;rft.jtitle=Journal+of+the+Optical+Society+of+America&amp;rft.pages=1297-1300&amp;rft.volume=70" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-25"><span class="mw-cite-backlink"><a href="#cite_ref-25">↑</a></span> <span class="reference-text"><a rel="nofollow" class="external text" href="https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf"><i>ImageNet Classification with Deep Convolutional Neural Networks, A. Krizhevsky, I. Sutskever and G. E. Hinton</i></a> (PDF; 1,4&nbsp;MB)</span>
</li>
<li id="cite_note-26"><span class="mw-cite-backlink"><a href="#cite_ref-26">↑</a></span> <span class="reference-text"><a rel="nofollow" class="external text" href="https://papers.cnl.salk.edu/PDFs/The%20The%20_'Independent%20Components_'%20of%20natural%20scenes%20are%20edge%20filters%201997-4064.pdf"><i>The “Independent Components” of Scenes are Edge Filters</i></a> (PDF; 1,3&nbsp;MB), A. Bell, T. Sejnowski, 1997, abgerufen am 17. November 2016.</span>
</li>
<li id="cite_note-27"><span class="mw-cite-backlink"><a href="#cite_ref-27">↑</a></span> <span class="reference-text">baeldung.com: <a rel="nofollow" class="external text" href="https://www.baeldung.com/cs/ai-convolutional-neural-networks"><i>Introduction to Convolutional Neural Networks</i></a></span>
</li>
<li id="cite_note-28"><span class="mw-cite-backlink"><a href="#cite_ref-28">↑</a></span> <span class="reference-text"><a rel="nofollow" class="external text" href="http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf"><i>ImageNet Classification with Deep Convolutional Neural Networks</i></a> (PDF; 1,4&nbsp;MB)</span>
</li>
<li id="cite_note-29"><span class="mw-cite-backlink"><a href="#cite_ref-29">↑</a></span> <span class="reference-text"><a rel="nofollow" class="external text" href="http://image-net.org/challenges/LSVRC/2016/">ILSVRC 2016 <i>Results</i></a></span>
</li>
<li id="cite_note-mcdns-30"><span class="mw-cite-backlink"><a href="#cite_ref-mcdns_30-0">↑</a></span> <span class="reference-text">Dan Ciresan, Ueli Meier, Jürgen Schmidhuber: <cite style="font-style:italic">Multi-column deep neural networks for image classification</cite>. In: <cite style="font-style:italic">2012 IEEE Conference on Computer Vision and Pattern Recognition</cite>. <i><a href="Institute_of_Electrical_and_Electronics_Engineers" title="Institute of Electrical and Electronics Engineers">Institute of Electrical and Electronics Engineers</a></i> (IEEE), New York, NY Juni 2012, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>3642–3649</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1109/CVPR.2012.6248110">10.1109/CVPR.2012.6248110</a></span>, <a href="ArXiv" title="ArXiv">arxiv</a>:<a rel="nofollow" class="external text" href="https://arxiv.org/abs/1202.2745v1">1202.2745v1</a> (<a rel="nofollow" class="external text" href="http://ieeexplore.ieee.org/xpl/articleDetails.jsp?arnumber=6248110">Online</a> [abgerufen am 9.&nbsp;Dezember 2013]).<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Abook&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Multi-column+deep+neural+networks+for+image+classification&amp;rft.au=Dan+Ciresan%2C+Ueli+Meier%2C+J%C3%BCrgen+Schmidhuber&amp;rft.btitle=2012+IEEE+Conference+on+Computer+Vision+and+Pattern+Recognition&amp;rft.date=2012-06&amp;rft.doi=10.1109%2FCVPR.2012.6248110&amp;rft.genre=book&amp;rft.pages=3642-3649&amp;rft.place=New+York%2C+NY&amp;rft.pub=Institute+of+Electrical+and+Electronics+Engineers+%28IEEE%29" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-31"><span class="mw-cite-backlink"><a href="#cite_ref-31">↑</a></span> <span class="reference-text"><i><a rel="nofollow" class="external text" href="https://ieeexplore.ieee.org/abstract/document/6835990">Improving multiview face detection with multi-task deep convolutional neural networks</a></i></span>
</li>
<li id="cite_note-32"><span class="mw-cite-backlink"><a href="#cite_ref-32">↑</a></span> <span class="reference-text"><span class="cite"><a rel="nofollow" class="external text" href="http://www.aclweb.org/anthology/W14-2405"><i>A Deep Architecture for Semantic Parsing.</i></a><span class="Abrufdatum"> Abgerufen am 17.&nbsp;November 2016</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=A+Deep+Architecture+for+Semantic+Parsing&amp;rft.description=A+Deep+Architecture+for+Semantic+Parsing&amp;rft.identifier=http%3A%2F%2Fwww.aclweb.org%2Fanthology%2FW14-2405&amp;rft.language=en">&nbsp;</span></span>
</li>
<li id="cite_note-33"><span class="mw-cite-backlink"><a href="#cite_ref-33">↑</a></span> <span class="reference-text"><span class="cite"><a rel="nofollow" class="external text" href="http://research.microsoft.com/apps/pubs/default.aspx?id=214617"><i>Learning Semantic Representations Using Convolutional Neural Networks for Web Search – Microsoft Research.</i></a> In: <i>research.microsoft.com.</i><span class="Abrufdatum"> Abgerufen am 17.&nbsp;November 2016</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=Learning+Semantic+Representations+Using+Convolutional+Neural+Networks+for+Web+Search+%E2%80%93+Microsoft+Research&amp;rft.description=Learning+Semantic+Representations+Using+Convolutional+Neural+Networks+for+Web+Search+%E2%80%93+Microsoft+Research&amp;rft.identifier=&amp;rft.date=&amp;rft.language=en">&nbsp;</span></span>
</li>
<li id="cite_note-34"><span class="mw-cite-backlink"><a href="#cite_ref-34">↑</a></span> <span class="reference-text"><span class="cite"><a rel="nofollow" class="external text" href="http://www.aclweb.org/anthology/P14-1062"><i>A Convolutional Neural Network for Modelling Sentences.</i></a> 17.&nbsp;November 2016<span style="display:none">;</span><span class="Abrufdatum" style="display:none"> abgerufen im 1.&nbsp;Januar 1</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=A+Convolutional+Neural+Network+for+Modelling+Sentences&amp;rft.description=A+Convolutional+Neural+Network+for+Modelling+Sentences&amp;rft.identifier=&amp;rft.date=2016-11-17&amp;rft.language=en">&nbsp;</span></span>
</li>
<li id="cite_note-35"><span class="mw-cite-backlink"><a href="#cite_ref-35">↑</a></span> <span class="reference-text"><span class="cite"><a rel="nofollow" class="external text" href="http://www.aclweb.org/anthology/D14-1181"><i>Convolutional Neural Networks for Sentence Classification.</i></a><span class="Abrufdatum"> Abgerufen am 17.&nbsp;November 2016</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=Convolutional+Neural+Networks+for+Sentence+Classification&amp;rft.description=Convolutional+Neural+Networks+for+Sentence+Classification&amp;rft.identifier=&amp;rft.date=&amp;rft.language=en">&nbsp;</span></span>
</li>
<li id="cite_note-36"><span class="mw-cite-backlink"><a href="#cite_ref-36">↑</a></span> <span class="reference-text"><span class="cite"><a rel="nofollow" class="external text" href="https://arxiv.org/pdf/1103.0398v1.pdf"><i>Natural Language Processing (almost) from Scratch.</i></a><span class="Abrufdatum"> Abgerufen am 17.&nbsp;November 2016</span> (englisch).</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=Natural+Language+Processing+%28almost%29+from+Scratch&amp;rft.description=Natural+Language+Processing+%28almost%29+from+Scratch&amp;rft.identifier=&amp;rft.date=&amp;rft.language=en">&nbsp;</span></span>
</li>
<li id="cite_note-37"><span class="mw-cite-backlink"><a href="#cite_ref-37">↑</a></span> <span class="reference-text"><span class="cite">heise online: <a rel="nofollow" class="external text" href="https://www.heise.de/newsticker/meldung/Maschinelle-Uebersetzer-DeepL-macht-Google-Translate-Konkurrenz-3813882.html"><i>Maschinelle Übersetzer: DeepL macht Google Translate Konkurrenz.</i></a> 29.&nbsp;August 2017,<span class="Abrufdatum"> abgerufen am 18.&nbsp;September 2017</span>.</span><span style="display: none;" class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&amp;rfr_id=info%3Asid%2Fde.wikipedia.org%3AConvolutional+Neural+Network&amp;rft.title=Maschinelle+%C3%9Cbersetzer%3A+DeepL+macht+Google+Translate+Konkurrenz&amp;rft.description=Maschinelle+%C3%9Cbersetzer%3A+DeepL+macht+Google+Translate+Konkurrenz&amp;rft.identifier=https%3A%2F%2Fwww.heise.de%2Fnewsticker%2Fmeldung%2FMaschinelle-Uebersetzer-DeepL-macht-Google-Translate-Konkurrenz-3813882.html&amp;rft.creator=heise+online&amp;rft.date=2017-08-29&amp;rft.language=de">&nbsp;</span></span>
</li>
<li id="cite_note-38"><span class="mw-cite-backlink"><a href="#cite_ref-38">↑</a></span> <span class="reference-text">Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness: <cite style="font-style:italic">Human-level control through deep reinforcement learning</cite>. In: <cite style="font-style:italic">Nature</cite>. <span style="white-space:nowrap">Band<span style="display:inline-block;width:.2em">&nbsp;</span>518</span>, <span style="white-space:nowrap">Nr.<span style="display:inline-block;width:.2em">&nbsp;</span>7540</span>, Februar 2015, <a href="Internationale_Standardnummer_f%C3%BCr_fortlaufende_Sammelwerke" title="Internationale Standardnummer für fortlaufende Sammelwerke">ISSN</a>&nbsp;<span style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://zdb-katalog.de/list.xhtml?t=iss%3D%220028-0836%22&amp;key=cql">0028-0836</a></span>, <span style="white-space:nowrap">S.<span style="display:inline-block;width:.2em">&nbsp;</span>529–533</span>, <a href="Digital_Object_Identifier" title="Digital Object Identifier">doi</a>:<span class="uri-handle" style="white-space:nowrap"><a rel="nofollow" class="external text" href="https://doi.org/10.1038/nature14236">10.1038/nature14236</a></span>.<span class="Z3988" title="ctx_ver=Z39.88-2004&amp;rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Ajournal&amp;rfr_id=info:sid/de.wikipedia.org:Convolutional+Neural+Network&amp;rft.atitle=Human-level+control+through+deep+reinforcement+learning&amp;rft.au=Volodymyr+Mnih%2C+Koray+Kavukcuoglu%2C+David+Silver%2C+...&amp;rft.date=2015-02&amp;rft.doi=10.1038%2Fnature14236&amp;rft.genre=journal&amp;rft.issn=0028-0836&amp;rft.issue=7540&amp;rft.jtitle=Nature&amp;rft.pages=529-533&amp;rft.volume=518" style="display:none">&nbsp;</span></span>
</li>
<li id="cite_note-rani2022-39"><span class="mw-cite-backlink"><a href="#cite_ref-rani2022_39-0">↑</a></span> <span class="reference-text">Rani, G. E., Venkatesh, E., Balaji, K., Yugandher, B., Kumar, A. N., &amp; SakthiMohan, M. (2022, April). An automated prediction of crop and fertilizer disease using Convolutional Neural Networks (CNN). In 2022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering (ICACITE) (pp. 1990–1993). IEEE.</span>
</li>
<li id="cite_note-sethy2021-40"><span class="mw-cite-backlink"><a href="#cite_ref-sethy2021_40-0">↑</a></span> <span class="reference-text">Sethy, P. K., Barpanda, N. K., Rath, A. K., &amp; Behera, S. K. (2020). Nitrogen deficiency prediction of rice crop based on convolutional neural network. Journal of Ambient Intelligence and Humanized Computing, 11(11), 5703-5711.</span>
</li>
</ol></div><!--htdig_noindex--><div><div class="zim-footer">
Dieser Artikel wurde von <a class="external text" title="Zuletzt bearbeitet am 2025-12-08" href="https://de.wikipedia.org/wiki/?title=Convolutional_Neural_Network&amp;oldid=262253469">Wikipedia</a> herausgegeben. Der Text ist unter <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.de">Creative Commons Attribution-Share Alike 4.0</a> verfügbar, sofern nicht anders angegeben. Für die Mediendateien können zusätzliche Bedingungen gelten.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
<script src="./_webp_/webpHandler.js"></script>

</body></html>